Back

Journal of the American Society for Mass Spectrometry

American Chemical Society (ACS)

Preprints posted in the last 30 days, ranked by how well they match Journal of the American Society for Mass Spectrometry's content profile, based on 37 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Orbitrap Collision Cross Section Measurements Enhance Isomer Annotations in Lipidomics

Ni, Z.; Ayzikov, K.; Makarov, A. A.; Moore, S.; Gaul, D. A.; Fort, K. L.; Fernandez, F.

2026-07-04 biochemistry 10.64898/2026.07.03.735735 medRxiv
Top 0.1%
10.2%
Show abstract

Despite advances in high-resolution mass spectrometry (HRMS), confident lipid annotation remains challenging due to the extensive chemical diversity of the lipidome and the prevalence of isomeric species. Ion mobility collision cross section (CCS) measurements provide structural information that complements HRMS; however, not all HRMS platforms can perform these measurements, necessitating a trade-off among mass resolution, accuracy, and robustness. Here, we introduce a method to infer lipid CCS values directly from liquid chromatography (LC)-Orbitrap MS experiments (Orbi). We show that Orbitrap mass analyzer pressure readings, and therefore CCS values, are influenced by the LC gradient solvent composition, requiring correction using isotopically labeled internal standards injected post-column. We also show that hundreds of lipid features can be assigned OrbiCCS values in a single LC run, with average precision better than 1% and an accuracy of 1-2% relative to reference DTCCS and TIMSCCS values. This excellent CCS accuracy not only enables more reliable annotation of lipid species in complex mixtures by matching OrbiCCS values to reference databases but also accelerates lipid structural elucidation based on the unknown's position in Orbi-retention time-m/z space.

2
RNabel-A Standalone Software Tool for Annotating Tandem Mass Spectra of Modified Ribonucleic Acids

Song, G.; Du, Y.-J. N.; Sun, R.; Dong, M.-Q.

2026-06-24 bioinformatics 10.64898/2026.06.22.733900 medRxiv
Top 0.1%
8.4%
Show abstract

Ribonucleic acid (RNA) modifications, with over 170 identified types, play diverse roles in cellular processes. The past decade has witnessed surging demand for accurate identification and localization of RNA modifications in both endogenous and synthetic therapeutic RNAs. With accurate spectral annotation for RNA, tandem mass spectrometry (MS/MS) can meet this demand. Here we present RNabel, a user-friendly software tool for in-depth annotation of MS/MS spectra of RNA oligonucleotides. RNabel considers a full set of backbone-cleavage ions (a, b, c, d, a-B, w, x, y, z) in which the ribonucleotide unit could be A, U, C, G, Y (pseudouridine), or I (Inosine). Additionally, RNabel considers 196 modifications on the base, the phosphoribose linkage, the 5' or the 3' terminus, or detachment of a sub-nucleotide fragment as a neutral or charged group. Users can create new components if needed, including ribonucleotides, modifications, neutral or charged groups that could detach from a ribonucleotide. RNabel efficiently processes large datasets in four acceptable formats including .mgf, .raw, .txt from msConvert, and RNabel batch files. Multiple statistical metrics are provided for quality assessment of spectral annotation. To accelerate RNA modification analysis, RNabel is made freely available for Mac and Windows users at https://github.com/songge1111/RNabel/releases. Graphic Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=116 SRC="FIGDIR/small/733900v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@8ccae5org.highwire.dtl.DTLVardef@15c8cfaorg.highwire.dtl.DTLVardef@12b93a2org.highwire.dtl.DTLVardef@1e9aab9_HPS_FORMAT_FIGEXP M_FIG C_FIG

3
Performance Evaluation of a Quantitative Metabolomics Workflow Incorporating Microchip Capillary Electrophoresis, Indexed Migration Time, and Single-Point External Calibration

Mellors, S.; Moss, C.; Redman, E. A.; Shuford, C.; Campbell, J. P.; Ramsey, J. M.; Coon, J.; Thompson, W.

2026-07-13 molecular biology 10.64898/2026.07.10.737294 medRxiv
Top 0.1%
7.6%
Show abstract

Capillary electrophoresis-mass spectrometry (CE-MS) offers unique analytical advantages for polar metabolite profiling but has remained underutilized in metabolomics relative to liquid chromatography-MS (LC-MS), in part due to challenges in managing migration time drift during data analysis. Here we introduce the use of indexed migration time (iMT) for easily managing this aspect of CE-MS data for metabolomics. Migration time indexing using a panel of stable isotope-labeled (SIL) amino acid reference standards, stored as an iRT database in Skyline, outperformed both uncorrected migration time and relative migration time (RMT) correction across three independent analytical batches spanning 90 samples from four biological matrices. The indexed migration time approach achieved sub-1% relative standard deviation (RSD) in migration index across batches, compared to up to [~]15% RSD for uncorrected migration times. Additionally, we evaluate the use of single-point external calibration in Skyline for the purposes of metabolite quantification from complex matrices in order to ease the burden of translational metabolite quantification from metabolomics using high-resolution mass spectrometry (HRMS). Single-point external calibration using a biological matrix-based calibrator was benchmarked against a 13-point linear calibration curve across a panel of amino acids; above 1 M, greater than 95% of back-calculated concentrations fell within {+/-}20% of multi-point calibration. Application of the complete workflow to plasma, serum, urine, and NIST Standard Reference Material (SRM)-1950 demonstrated low inter-batch variability by principal components analysis, broad metabolite coverage across 126 quantifiable analytes, and strong quantitative concordance (Deming slope = 0.862, pseudo-R2 = 0.994, n = 64 analytes) with an independent comprehensive reference dataset for NIST SRM-1950. Together, these results establish a practical mCE-HRMS metabolomics workflow that bridges targeted and discovery metabolomics paradigms and lays the groundwork for single-point external calibration as a powerful tool for translational metabolomics.

4
DESI-MS-Based Analysis of Drug Distribution in Human Renal Cystic Tissue Using the Chorioallantoic Membrane (CAM) as a 3D In Vivo Model

Dettmer, K.; Hehemann, A. M. E.; Schueler, J.; Heckscher, S.; Gross, V.; May, M.; Nuebel, B.; Wullich, B.; Buchholz, B.; Werner, J. M.; Jantsch, J.; Gronwald, W.; Takats, Z.; Oefner, P. J.; Schmidt, K. M.; Haerteis, S.

2026-07-01 biochemistry 10.64898/2026.07.01.735776 medRxiv
Top 0.1%
5.6%
Show abstract

The chorioallantoic membrane (CAM) model represents a promising three-dimensional in vivo platform for preclinical drug testing in human tissues. In this study, we investigated whether the tissue penetration and distribution of benzbromarone, a known inhibitor of the Ca2+ activated chloride channel TMEM16A and potential therapeutic agent for autosomal dominant polycystic kidney disease (ADPKD), can be successfully visualized in human renal cyst tissue cultured on the CAM. To this end, desorption electrospray ionization mass spectrometry imaging (DESI-MSI) combined with an ultrahigh-resolution time-of-flight mass spectrometer was employed. We achieved spatially resolved molecular mapping of endogenous metabolites and lipids as well as the applied compound. MSI enabled clear differentiation between CAM and cystic tissue based on their distinct lipid profiles. Benzbromarone was reproducibly detected in the cyst specimens and exhibited selective accumulation along the cyst epithelium, which is considered the principal site of action. These observations were complemented by multivariate analyses including Uniform Manifold Approximation and Projection (UMAP), and sparse multinomial logistic zero-sum classification. The data-driven approach confirmed molecular differences between tissue types and allowed accurate classification of drug-treated and untreated regions. This study demonstrates that topically applied benzbromarone penetrates human renal cyst tissue in the CAM model and localizes to pharmacologically relevant tissue regions, notably the location of the Ca2+ activated chloride channel TMEM16A in the epithelial lining. The integration of high-resolution DESI-MSI with advanced statistical analysis provides a robust and label-free method to study drug distribution in human tissue grafts. Our findings contribute to the advancement of translational research in analytical chemistry and pharmacology.

5
Enhanced proteome relative quantification using refined quantotypic spectral libraries

Barnes, B. A.; Alharbi, H.; Unwin, R.

2026-07-10 bioinformatics 10.64898/2026.07.06.736793 medRxiv
Top 0.1%
5.5%
Show abstract

Plasma proteomics is used for a variety of applications including biomarker discovery, disease monitoring, and drug development. Data-independent acquisition (DIA) has vastly improved the breadth of proteins that are identified from samples; however, given challenges in reproducibility and translation, it is critical that the quantitative performance of these methods is reliable. Analysis of global proteomics data typically incorporates information from all detected peptides. However, some peptides do not reflect their parent protein amount, due to irreproducible digestion, modification, analytical interferences or instability. We hypothesise that including these peptides impacts protein relative quantification, and thus, a refined spectral library containing only quantitatively representative peptides provides superior protein quantification. By analysing a defined multi-species spike-in model, we show that refining a plasma spectral library by removing precursors that fail to meet quality control metrics (25.4% of all identified precursors) reduces noise and variability, improving precision, accuracy and differential abundance analysis by up to [~]11%, with minimal identification losses and substantial reduction in computational demand. This demonstrates proof-of-concept that refining spectral libraries produces results that prioritize quantification quality over quantity. This approach could enable development of universal tissue-specific refined spectral libraries able to improve quantification quality with easy implementation and minimal processing time. Significance of the StudyAs DIA mass spectrometry proteome depth increases, the quality of the associated protein quantifications must be considered alongside identification breadth, particularly in complex matrices such as plasma, which presents additional technical challenges. The spectral library used for protein identification and quantification is a critical determinant of DIA performance, and its composition requires considerable consideration. This work illustrates an initial step toward improving protein quantification starting at the spectral library level by filtering precursors which are poor quantitative representatives of their parent proteins. In doing so, the resulting data is more reliable for downstream and biological interpretation, with fewer false differential abundance assignments and reduced quantitative noise. As such, this work represents a broader shift away from the habitual focus of MS workflows on maximising the number of protein and differential abundance identifications and instead prioritises the quality of quantification over quantity. These initial findings lay the groundwork for further development of spectral library refinement strategies, with the potential to continue improving the accuracy and precision of protein quantification in DIA-based proteomics.

6
MassSpectrum Analyzer: An interactive platform for proteomic searching parameter refinement and peptide modification focused re-scoring

Karlic, K. I.; Scott, N. E.

2026-06-28 bioinformatics 10.64898/2026.06.22.733873 medRxiv
Top 0.1%
4.3%
Show abstract

Peptide spectrum annotation is critical for the assignment of peptides and the localisation of modifications. While many existing tools provide spectrum annotation capacities, they often lack the flexibility required to allow bespoke spectral annotation of peptides containing multiple labile modifications or the accurate assignment of peptides in which fragmentation deviates from canonical patterns. In these cases, user-guided annotation is widely used to improve assignment completeness, however it typically does not integrate peptide scoring, making it challenging to assess the empirical improvement of the associated annotation and its impact on downstream false-discovery rate estimations. Here, we introduce an interactive annotation environment, the 'MassSpectrum Analyzer', which aims to streamline the exploration and analysis of modified peptides by enabling user-defined customisation with peptide scoring. Using (2-Aminoethyl)trimethylammonium carboxyl-derivatised peptides and glycopeptides as case studies we demonstrate the capacity of the MassSpectrum Analyzer to rapidly explore and allow the assessment of modified peptide datasets. By enabling direct assessment of the impact of user-guided choices on peptide scoring, we show how the detection of highly modified peptides can be improved through post-search integration of modification fragmentation information in a statistically robust manner. Similarly, by permitting comparisons of peptide ion intensities across spectra, we show that global fragmentation patterns can be quantified allowing the interrogation of trends that only become clear when spectra are assessed en masse. Combined, the MassSpectrum Analyzer streamlines the generation of publication-ready spectra and provides a means to assess how the inclusion of annotated features influences assignment scores.

7
Reducing background ion burden in tributylamine ion-pairing LC-MS improves signal intensity and feature coverage in metabolomics

Tarach, A. R.; Vincent, M. P.; Ellis, A. E.; Isaguirre, C. N.; Caudy, A. A.; Sheldon, R. D.

2026-06-25 biochemistry 10.64898/2026.06.24.734057 medRxiv
Top 0.1%
4.3%
Show abstract

Background chemical ions are a pervasive but often underappreciated limitation in LC-MS metabolomics, where they can suppress analyte signal, obscure endogenous metabolites, increase spectral complexity, and consume MS/MS acquisition events. Tributylamine (TBA) ion-pairing reversed-phase LC-MS provides stable retention and broad coverage of polar anionic metabolites, including central carbon intermediates, nucleotides, cofactors, and bile acids, but the back-ground burden introduced by the ion-pairing reagent itself has not been systematically addressed. Here, we identify commercial TBA as a major source of nonbiological contaminant ions and develop a practical strategy to reduce background burden while preserving metabolite coverage. Serial solid-phase extraction of TBA using orthogonal reversed-phase, strong anion-exchange, and strong cation-exchange sorbents removed chemically diverse contaminants, including isobaric background ions that interfered with endogenous hydroxybutyrate isomers. We further optimized the workflow by reducing medronic acid concentration, restricting medronic acid to the organic mobile phase, replacing phosphoric-acid column conditioning with metal-passivated column hardware, and adding EDTA to the sample reconstitution solvent to improve citrate detection. In mouse liver extracts, the optimized method increased signal intensity for most annotated metabolites and improved the fraction of full-scan ion current attributable to target analytes. Method optimization also altered compound-specific retention behavior, resolving some co-elution-based interferences while introducing new suppression relationships for selected analytes. Across mouse liver, human B lymphocytes, and NIST SRM 1950 plasma, the optimized workflow increased total feature detection by 45%, 72%, and 42%, respectively, and improved the number of low-variance features, precursors with data-dependent MS/MS spectra, and MS/MS library matches. These findings establish background-ion mitigation as a central design principle for LC-MS method development. More broadly, this work provides a generalizable framework for identifying, reducing, and validating reagent- and additive-derived background to improve targeted and untargeted LC-MS data quality.

8
Agentic AI for Structural Elucidation and Discovery of Drug Metabolites from Mass Spectrometry Data

Wang, X.; Patan, A.; Zhao, H. N.; Charron-Lamoureux, V.; Shin, Y.; Petras, D.; Hong, Y.; Bowen, B. P.; Northen, T. R.; Dorrestein, P. C.; Wang, M.

2026-06-26 bioinformatics 10.64898/2026.06.23.734138 medRxiv
Top 0.2%
3.9%
Show abstract

The majority of chemical signals detected in public metabolomics repositories remain structurally undefined. Large language models (LLMs) are probabilistic systems whose capacity to generate outputs beyond their training data, which can cause hallucinations, makes them also potentially suited to hypothesize structures for molecules that have never been described. We aimed to build a system that could harness this LLM generative capacity combined with domain specific tools/framework to constrain hallucination and produce validated discoveries. We developed a GNPS2 agentic AI system that interprets LC-MS/MS data by integrating spectral alignment, molecular formula inference, rule-based structural enumeration, machine learning-based spectrum prediction, and translates natural language hypotheses from domain experts into dynamically generated analytical workflows. We demonstrate the annotation of unknown drug metabolites from public data guided by chemical hypotheses. The agent predicted, and we experimentally confirmed, a phosphorylated hydroxyzine, an acetaminophen-p-coumaric acid ester, and identified two new oxidative ibuprofen-carnitine conjugates from public repositories. These results demonstrate that LLM-driven agentic reasoning, when combined with domain expertise, can indeed generate experimentally testable structural hypotheses for previously uncharacterized metabolites leveraging pan repository data.

9
A Raman Spectroscopy-Based Method for Label-Free Discrimination of Human Inhibin α, Inhibin B, and Activin A

Xiao, W.; Dai, Y.; Martinez Gallardo Quijano, S.; Tsigkou, A.; Kotsifaki, D.

2026-07-06 biochemistry 10.64898/2026.07.04.735879 medRxiv
Top 0.2%
3.6%
Show abstract

Members of the transforming growth factor-{beta} (TGF-{beta}) superfamily, including inhibins and activins, are structurally related glycoprotein dimers that regulate reproductive and endocrine signaling. Their high degree of molecular similarity presents challenges for label-free analytical discrimination. To evaluate the ability of Raman spectroscopy to distinguish closely related TGF-{beta} superfamily proteins based on intrinsic vibrational fingerprints. Raman spectra of recombinant human Inhibin -subunit, Inhibin B ({beta}B homodimer), and Activin A ({beta}A--{beta}A) were acquired using confocal Raman microscopy with 532 nm excitation. Spectra were baseline-corrected, area-normalized, and analysed using principal component analysis (PCA). Distinct spectral signatures were observed across the 500--1800 cm-1 region. Differences within the S--S stretching region (500--550 cm-1) were consistent with variations in disulfide-bond environments, with the Inhibin -subunit exhibiting the highest relative intensity in this region. Variations in the amide I band (1600--1700 cm-1) suggested differences in protein secondary structure, while aromatic amino acid vibrations provided additional discriminatory features. PCA revealed clear clustering and separation of all three protein classes based on their Raman fingerprints. Raman spectroscopy enables label-free differentiation of structurally related endocrine glycoproteins and demonstrates potential for the structural characterization and classification of inhibin and activin proteins within the TGF-{beta} superfamily.

10
MetaboCensoR: A Shiny Application for Data Filtering in Untargeted LC-MS Metabolomics to Enhance Interpretability

Plyushchenko, I. V.; Luzzatto-Knaan, T.

2026-07-07 bioinformatics 10.64898/2026.07.02.735197 medRxiv
Top 0.2%
3.5%
Show abstract

Untargeted LC-MS metabolomics datasets often contain large numbers of redundant and non-informative features arising from background contaminants, multiple ion forms, poorly integrated peaks, and other low-quality signals. These features complicate downstream analysis by inflating feature space, degrading molecular networks, impeding pathway analysis, and obscuring statistically meaningful changes. Here, we present MetaboCensoR, an input-versatile Shiny application and local R package for analyte-centric peak table filtering. The workflow integrates four complementary modules for blank filtering, redundant ion-species filtering, quality-control filtering, and peak-based filtering. MetaboCensoR also provides interactive threshold optimization, exportable annotation tables, and synchronized filtering of associated .mgf files. The approach was evaluated across three independent datasets covering plant extracts, human cell lines, and bacterial interactions. Across these case studies, data filtering reduced feature redundancy and improved downstream interpretation in feature-based molecular networking, pathway-level functional analysis, and differential abundance testing, while preserving known target metabolites. These results show that systematic peak table filtering can substantially improve the interpretability and analytical value of untargeted metabolomics data.

11
Application of class-balancing algorithms to diverse plasma metabolomics datasets using brain tumor as an example

Godlewski, A.; Solowiej, K.; Mojsak, P.; Godzien, J.; Zelkowska, J.; Kretowski, A.; Lyson, T.; Burdukiewicz, M.; Kaminski, K.; Ciborowski, M.

2026-07-07 bioinformatics 10.64898/2026.07.02.735756 medRxiv
Top 0.2%
3.4%
Show abstract

Class imbalance remains a challenge in metabolomics research, where biological and technical variability can affect statistical inference and machine learning (ML) performance. Class-balancing algorithms address this issue by either increasing minority-class observations or reducing the number of majority-class samples. This study evaluated the impact of oversampling and undersampling algorithms on targeted and untargeted metabolomics datasets derived from LC-MS and GC-MS analyses of plasma samples from patients with glioblastoma, meningioma, and controls. Synthetic Minority Oversampling Technique (SMOTE) and Random Undersampling (RUS) were applied to balance the datasets, and their effects on data distribution, inter-feature correlations, and machine learning model performance were compared. RUS preserved the original feature distributions but reduced representativeness by removing the majority-class samples. In contrast, SMOTE introduced synthetic samples that altered covariance structures, increasing the risk of overfitting, particularly in small datasets (n=10). These effects diminished with larger groups (n=30), partially restoring correlations between metabolites. Model performance varied across the class-balancing algorithms. Random Forest classifiers benefited from both balancing methods, with undersampling often yielding higher F1 scores, whereas Support Vector Machine models showed reduced classification performance. These findings highlight the importance of selecting class-balancing strategies based on dataset size, analytical platform, and ML algorithm in metabolomics studies.

12
Data Independent Acquisition Pipeline for Microbiome Samples (Microbe-DIA)

Obermiller, S. A.; Lipton, M. S.; Piehowski, P. D.; Bilbao, A.; McCue, L. A.; Prozapas, V. N.; Attah, I. K.

2026-07-14 microbiology 10.64898/2026.07.13.738261 medRxiv
Top 0.2%
3.1%
Show abstract

The functional complexity inherent in microbiomes complicates analytical approaches aimed at defining phenotype. As proteins are the functional effectors of microbiome phenotypes, improving the performance of mass spectrometry-based metaproteomics is critical to achieving the functional characterization of these systems. Data-independent acquisition (DIA) improves protein coverage and reduces data missingness when compared to data-dependent acquisition (DDA) in metaproteomics. However, the application of DIA to complex microbial systems remains constrained by analytical throughput and computational scalability. Here, we optimized LC-MS/MS acquisition parameters for both DDA and DIA using a model microbiome, demonstrating how DIA enables increased sample throughput without compromising quantitative performance. In addition, we demonstrated a computationally efficient, library-free DIA workflow that overcomes reliance on empirical spectral libraries. Our analytical and computational innovations establish a scalable and cost-effective pipeline for metaproteomics of complex microbial communities.

13
Development of a Matrix-Matched Calibration Curve for Multi-Site Quantification of Neu5Gc-Bearing N-Glycans

DeBono, N. J.; Moh, E. S.; Poole, J.; Packer, N. H.; Day, C. J.; Jennings, M. P.; Kolarich, D.; Ashwood, C.

2026-07-15 biochemistry 10.64898/2026.07.14.738351 medRxiv
Top 0.2%
3.1%
Show abstract

N-glycolylneuraminic acid (Neu5Gc) has been repeatedly associated with human cancer, but reliable detection has remained elusive, generating controversy regarding its presence in human samples. To address this, matrix-matched calibration curves, which have been pioneered in proteomics and metabolomics for assessing changes in complex mixtures, were measured of released N-glycans at four orders of magnitude dynamic range in defined mixtures, systematically benchmarking Neu5Gc-containing N-glycan detection across multiple LC-MS platforms and sites. Orthogonally, the gold-standard analytical method, consisting of fluorescence detection of labelled monosaccharides separated by LC, was applied to the same samples, yielding absolute concentrations of Neu5Gc. LC-MS demonstrated an extended detection range of three or more orders of magnitude while retaining intact N-glycan measurement, improving assay specificity and enabling detection of the variety of Neu5Gc-bearing N-glycans. By combining orthogonal dimensions of evidence, including chromatographic separation, isotopic distribution matching, and composition-confirming MS/MS, LC-MS confidently resolved Neu5Gc signals from noise, even at low abundance. In comparison, DMB-LC-FLR was limited to two orders of magnitude dynamic range, insufficient for detection of Neu5Gc in commercially available pooled human sera. These findings strongly support that DMB-LC-FLR assay specificity and sensitivity are insufficient for Neu5Gc detection in human samples due to noise overwhelming the Neu5Gc signal. By establishing a reusable benchmarking framework for future glycomic studies, we aim to use LC-MS to improve the measurement of Neu5Gc in clinical samples.

14
A Pubic Hair Is 172 Times More Pubic Than a Scalp Hair

Ogata, N.; MATSUDA, T.

2026-07-01 bioengineering 10.64898/2026.06.25.734686 medRxiv
Top 0.2%
2.6%
Show abstract

Human hair is a common contaminant in GMP-controlled manufacturing environments, and its identification is important for contamination source investigation and corrective action. Because human hair can originate from multiple body sites, it is often necessary to determine not only the species of origin but also the anatomical source of the hair. Conventional forensic approaches distinguish scalp hair from body hair by microscopic examination of cuticle patterns, medullary structure, cross-sectional morphology, and pigment distribution. However, these methods depend on examiner expertise, are difficult to apply to damaged specimens, and provide limited quantitative information. In this study, we developed a proteomics-based approach for distinguishing scalp hair from pubic hair using identical sample preparation and analytical workflows. Comparative proteomic analysis identified keratin-associated proteins KAP 4-3 and KAP 9-6 as enriched in scalp hair, whereas cuticular keratins Ha7 and Ha8 were strongly enriched in pubic hair. Amino acid composition analysis further revealed that scalp hair-enriched proteins were highly cysteine-rich, consistent with sulfur-rich cross-linking matrix proteins, whereas pubic hair-enriched proteins exhibited characteristics of structural keratin filaments. These results demonstrate that proteomic signatures can provide a quantitative and objective means of determining the anatomical origin of human hair and may contribute to contamination source tracing in GMP manufacturing and forensic investigations.

15
Hydration and H/D exchange-dependent infrared signatures of the GCN4 leucine zipper

Bhuvanendran, H.; Brunner, C. M.; Kempf, H.; Moro, J. L.; Roubieu, E.; Turbant, F.; Mateus, A.; Lin, H.; Das, L.; Malyshev, D.; Johns, B.; Parracino, A.; Pastore, A.; Peters, J.; Cortajarena, A. L.; Zanetti Polzi, L.; Maccaferri, N.

2026-06-26 biochemistry 10.64898/2026.06.26.731617 medRxiv
Top 0.2%
2.4%
Show abstract

Attenuated total reflectance Fourier-transform infrared (ATR-FTIR) spectroscopy of proteins in aqueous solution is often limited by water absorption and other optical artifacts. To overcome these limitations, we evaluated the structural features and hydrogen-deuterium exchange (HDX) kinetics of the -helical protein GCN4 in both hydrated (wet) and vacuum-dried (dry) states. While solvent heavily mask the second-derivative spectra of wet samples, vacuum drying yielded a thin, protein-rich film on the ATR crystal, significantly enhancing the signal-to-noise ratio and resolving the protein features without altering the native structure. Dry-state analysis clearly resolved the Amide I, Amide II, and deuterium-shifted Amide II' (1450 cm-1) bands. Notably, second-derivative analysis of the dry spectra of the HDX samples revealed a bimodal Amide I distribution consisting of a stationary band at 1653 cm-1 from the solvent-inaccessible regions and an isotopically sensitive band shifting from 1648 cm-1 to 1644 cm-1 from solvent-accessible regions. These results demonstrate that vacuum-dried ATR-FTIR spectroscopy effectively eliminates solvent masking, providing the spectral clarity required to resolve discrete -helical sub-populations after deuteration.

16
Incorporating Surfaced-Induced Dissociation Mass Spectrometry Data into an AlphaFold-derived deep learning network improves protein structure prediction

Bolz, R. M.; Day, E. H.; Drake, Z. C.; Harvey, S. R.; Wysocki, V. H.; Lindert, S.

2026-06-29 biochemistry 10.64898/2026.06.26.734850 medRxiv
Top 0.2%
2.3%
Show abstract

Surface-Induced Dissociation native Mass Spectrometry (SID-nMS) is a tandem MS activation method that yields information on the connectivity and stoichiometry of protein complexes. While insufficient for direct structure elucidation, the data derived from SID-nMS has considerable potential to inform multimeric protein structure prediction. We hypothesized that incorporating this data into a machine-learning framework could improve multimer prediction accuracy beyond that of existing deep-learning methods. To this end, we developed SIDFold, a novel AlphaFold-based deep-learning network. SIDFold is the first AlphaFold-like network to leverage experimental data during protein complex prediction, and the first deep-learning network to utilize nMS data for structure prediction. We benchmarked SIDFold on the BETA protein set, and observed an improvement in RMSD in 138 of 227 cases including 27 targets in which the predicted structure attained near-native accuracy. We then evaluated the network on 20 proteins with experimental SID-nMS data, yielding an improved RMSD in 18 cases, with five of these cases improving to a high-accuracy complex. Finally, we tested SIDFold against a previously published SID-guided Rosetta docking method, where we saw improvement in 13 of 16 proteins. SIDFold is freely available on GitHub, with example files and commands available in the Supplementary Information.

17
Protein Aggregation Capture for Top-down Proteomics

Feltenstein, I. G.; Drown, B. S.

2026-07-03 biochemistry 10.64898/2026.07.02.736076 medRxiv
Top 0.3%
2.1%
Show abstract

Proteins are dynamically regulated by a myriad of post-translational modifications (PTMs) that control their stability, conformation, activity, subcellular localization, and local interactions. Capturing the precise composition of these various modification states, or proteoforms, is a principal objective of top-down proteomics (TDP). By ionizing intact proteoforms and combining measurements of precursor ion and fragment ion masses, the position, stoichiometry, and combination of PTMs can be determined. Despite the highly valuable measurements that TDP can provide, it is typically less sensitive than corresponding peptide-level analysis with many reports utilizing input material in the microgram to milligram range. Contributing to this lack of sensitivity is the risk of sample loss due to non-specific binding to surfaces during sample preparation. The most widely employed sample preparation approaches for TDP either require high sample input (e.g. precipitation and ultra-filtration) or fail to effectively remove surfactants (e.g. solid-phase extraction). These limitations have hindered advancement of targeted TDP applications involving immunoprecipitation and other enrichment strategies. Bead-assisted protein aggregation, also referred to as single-pot, solid-phase-enhanced sample preparation (SP3), has emerged as a popular sample preparation strategy for bottom-up proteomic workflows, but has only been used in TDP with secondary ion exchange chromatography cleanup. We envisioned a magnetic bead based protein cleanup approach that proceeds directly to MS analysis with judicious choice of bead surface chemistry and elution conditions. Here we report a sample preparation method using hydroxyl-functionalized magnetic beads for top-down proteomics applications.

18
Development of an Ethylenediaminetetraacetic Acid-Enhanced Deep Proteomic Profiling Method for Dried Blood Spots and Its Application in Mouse Disease Models

Nakajima, D.; Kanno, T.; Okuda, Y.; Mitsui, H.; Konno, R.; Ueyama, N.; Endo, Y.; Ohara, O.; Kawashima, Y.

2026-07-14 molecular biology 10.64898/2026.07.13.738354 medRxiv
Top 0.3%
2.1%
Show abstract

Dried blood spots (DBS) are well-established microsamples used in clinical testing and newborn screening. However, their use in deep proteomics is hindered by highly abundant blood proteins and inefficient protein recovery from filter paper matrices. The non-targeted analysis of non-specifically DBS-absorbed proteins (NANDA) workflow partially overcomes the impact of abundant blood proteins and has enabled the identification of over 5,000 proteins from DBS samples. Nonetheless, residual abundant proteins, including hemoglobin and fibrinogen, constrain deep proteomic analysis. Therefore, this study aimed to evaluate the effects of the metal chelator ethylenediaminetetraacetic acid (EDTA) on the depth of DBS proteomic analysis. An optimized EDTA-enhanced NANDA protocol that incorporated a 100 mM EDTA wash step was compatible with standard DBS collection procedures and required no modification of current clinical workflows, markedly enhancing the depletion of abundant proteins and facilitating its potential use in clinical and translational settings. When combined with Orbitrap Astral data-independent acquisition mass spectrometry, this approach enabled the single-shot identification of more than 7,000 proteins from DBS samples; to the best of our knowledge, this represents the deepest proteome coverage reported to date, and the workflow further supported high-throughput and highly reproducible analyses. Additionally, its application to mouse disease models revealed disease-specific systemic immune signatures from minimal blood volumes. Collectively, these results establish EDTA-enhanced NANDA as a practical and scalable workflow that overcomes longstanding limitations of DBS proteomics, thereby enabling deep, high-throughput, minimally invasive proteomic profiling across diverse biological and experimental contexts.

19
In Vivo Quantification of Histone Acetylation Turnover and Acetyl-CoA Sources Using 2H2O Metabolic Labeling and High-Resolution Mass Spectrometry.

Arias-Alvarado, A.; Sabir, U.; Ilchenko, S.; Parrish, S.; Aghayev, M.; He, W.; Tsai, T.-H.; Zhang, G.; Kasumov, T.

2026-06-29 biochemistry 10.64898/2026.06.26.734905 medRxiv
Top 0.3%
2.1%
Show abstract

Dysregulated histone acetylation links cellular metabolism to gene expression, but measuring its in vivo turnover remains technically challenging. Here, we introduce a 2H2O-based metabolic labeling method coupled with high-resolution Orbitrap mass spectrometry to quantify in vivo histone acetylation dynamics. The approach leverages differing deuterium incorporation rates between fast-labeling acetyl groups and slow-labeling peptide backbones. A two-tier analytical workflow uses full-scan mass spectrometry for mono-acetylated peptides, combined with parallel reaction monitoring (PRM) to resolve site-specific turnover and stoichiometry. Furthermore, monitoring acetyl-group plateau 2H enrichment enables the evaluation of specific substrate contributions to the acetyl-CoA pool supporting histone acetylation. To demonstrate biological utility, we applied this approach to mice maintained on a high-carbohydrate diet or subjected to 48-h fasting to assess nutrient-dependent histone acetylation dynamics. Acetyl-group labeling reflected the metabolic origin of acetyl-CoA, showing greater 2H enrichment in the fed state and reduced enrichment during fasting due to increased utilization of unlabeled fatty acid-derived acetyl-CoA. Fasting accelerated acetylation turnover across multiple histone sites and reduced overall acetylation stoichiometry. Quantitative tracing revealed that fatty acid oxidation becomes an important contributor to histone acetylation during fasting, whereas glucose remains the predominant source of nucleo-cytosolic acetyl-CoA (supplying > 60% of acetylation used carbon). This approach enables simultaneous in vivo assessment of histone acetylation turnover, site occupancy, and acetyl-CoA substrate utilization, offering a robust platform to investigate metabolic-epigenetic crosstalk in health and disease.

20
Kinetic Lipidomics: Quantifying in vivo changes in lipid metabolism using metabolic labeling

Nielsen, C.; Denton, R.; Driggs, B.; Gates, S.; Hilton, T.; Naylor, B.; Quilling, C.; Virgin, K.; Cutler, K.; Sorensen, M.; Poulson, M.; Snedaker, P.; Hernandez, Z.; Transtrum, M.; Price, J. C.

2026-07-01 biochemistry 10.64898/2026.06.29.735310 medRxiv
Top 0.3%
2.1%
Show abstract

Lipid metabolism reflects the dynamic balance between metabolic turnover and concentration. Kinetic mass spectrometry (MS) enables direct quantification of molecular turnover in vivo. Previous work has shown that MS-based kinetic proteomics has provided powerful insights into proteome regulation. Analogous lipidome-wide kinetic measurements remain limited by challenges in defining molecule-specific labeling behavior. Here, we extend kinetic MS to untargeted lipidomics. Isotope labeling with deuterated water (2H2O) is commonly used for monitoring turnover of palmitate and other select lipids by measuring labeling of stable CH positions with deuterium (2H). Here, we extend the deuterium-incorporation model underlying these targeted lipid turnover assays to support untargeted analysis of all detectable lipids. This allows us to empirically quantify the effective fraction of endogenous synthesis (Asyn) and the turnover rate (k) across hundreds of lipid species simultaneously. One central barrier to lipidome-wide kinetic modeling is determining the endogenous number of deuterium-labeling sites for each molecule (nL) which is required to estimate Asyn and k accurately. The nL value is an essential component of biological kinetic assays. In kinetic proteomics, curated amino acid nL libraries enable peptide-level modeling by summing sequence-specific labeling-site values, but comparable resources are lacking for lipids and may not generalize across metabolic states or non-mammalian systems. Yet, gaps remain for lipids and for amino acids in modified metabolic conditions or non-mammalian biologies. Here, we empirically determine lipid nL values and validate the process with peptides against an nL library. To evaluate this strategy in a biologically relevant setting, we applied it to brain tissue from transgenic mice expressing human ApoE isoforms, where altered lipid transport and metabolism are implicated in Alzheimers disease risk. These data validate the method in a clinically relevant context and suggest that genotype-dependent metabolism can alter empirically determined lipid nL values.